Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Ravikumar V, Dr. S. Sasirekha
DOI Link: https://doi.org/10.22214/ijraset.2026.84701
Certificate: View Certificate
Diabetes prediction plays an important role in re-ducing long-term health risks by enabling early medical interven-tion. Although machine learning models have been widely applied to this task, many existing studies emphasise predictive accuracy while giving comparatively little attention to the reliability, interpretability, and stability of the resulting decisions. This paper develops a reliability-aware and interpretable machine learning framework for diabetes prediction from structured clinical data. Three complementary models—Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost)—are trained on the Pima Indians Diabetes dataset so that both simple linear and complex non-linear relationships are captured. Beyond conventional discrimination metrics, the reliability of the predicted probabilities is quantified using the Brier score and reliability (calibration) diagrams. Interpretability is addressed with SHapley Additive exPlanations (SHAP) at both the global (cohort) and local (individual patient) levels. Because different models frequently emphasise different predictors, we formalise a Feature Consistency Index (FCI) that quantifies the cross-model agreement of SHAP-derived feature importance and combines it with normalised importance into a single ranking score. Finally, a perturbation-based robustness analysis measures the sensitivity of each model’s output to small changes in the input record. Experi-mentally, XGBoost achieves the highest discrimination (accuracy 0.7597, ROC-AUC 0.8374), whereas Random Forest attains the best-calibrated probabilities (Brier score 0.1646), demonstrating that discrimination and reliability are not interchangeable. The FCI identifies Glucose and BMI as simultaneously the most influential and the most consistently attributed predictors, while Blood Pressure and Skin Thickness are both weak and unstable. Under a 5% Gaussian perturbation of a representative patient record, the linear and bagged models shift by less than 0.01 in predicted probability, whereas the boosted model shifts by 0.0386, revealing an accuracy–stability trade-off that a purely accuracy-driven evaluation would not expose.
The text presents a machine-learning framework for diabetes prediction that goes beyond conventional accuracy measures by evaluating models on four dimensions: accuracy, reliability, interpretability, and stability.
This paper presented a reliability-aware and interpretable machine learning framework for diabetes prediction from structured clinical data, evaluating Logistic Regression, Ran-dom Forest, and XGBoost along four coordinate axes: dis-crimination, probabilistic reliability, interpretability, and ro-bustness. The results show that these axes are genuinely distinct. XGBoost achieved the best discrimination (accuracy 0.7597, ROC-AUC 0.8374), Random Forest the best-calibrated proba-bilities (Brier score 0.1646) together with the highest recall on the diabetic class (0.76), and Logistic Regression and Random Forest were markedly more robust to input perturbation than XGBoost, whose predicted probability moved four to five times further under a 5% measurement-scale perturbation. Model selection therefore cannot be settled by accuracy alone, and the appropriate choice depends on how the model’s output will be consumed clinically. The proposed Feature Consistency Index quantified the degree to which SHAP attributions agree across model fam-ilies, converting a normally invisible source of explanatory uncertainty into a reported quantity. It identified Glucose and BMI as both the most influential and the most consistently attributed predictors; Age, the Diabetes Pedigree Function, and Pregnancies as a contested middle band whose internal ordering is not robust to model choice; and Blood Pressure and Skin Thickness as consistently uninformative. Patient-level analysis further showed that the cohort ranking does not transfer to individuals, supporting per-patient explanation in any deployed interface. Future work follows directly from the limitations of Sec-tion VI. The priorities are: external validation on an indepen-dent and demographically broader cohort; replacement of the single split with repeated stratified cross-validation reporting confidence intervals; a full Monte Carlo robustness study across the test set at multiple noise levels; scale-normalisation of the FCI so that its absolute values become interpretable, together with a comparison against rank-based agreement statistics such as Kendall’s W ; multiple imputation or explicit missingness indicators in place of median imputation; and the addition of post-hoc recalibration by Platt scaling or isotonic regression, which would allow discrimination and calibration to be optimised separately rather than traded off.
[1] International Diabetes Federation, IDF Diabetes Atlas, 10th ed. Brussels, Belgium: International Diabetes Federation, 2021. [2] I. Kavakiotis, O. Tsave, A. Salifoglou, N. Maglaveras, I. Vlahavas, and Chouvarda, “Machine learning and data mining methods in diabetes research,” Computational and Structural Biotechnology Journal, vol. 15, pp. 104–116, 2017. [3] Q. Zou, K. Qu, Y. Luo, D. Yin, Y. Ju, and H. Tang, “Predicting diabetes mellitus with machine learning techniques,” Frontiers in Genetics, vol. 9, art. 515, 2018. [4] M. Maniruzzaman, M. J. Rahman, B. Ahammed, and M. M. Abedin, “Classification and prediction of diabetes disease using machine learning paradigm,” Health Information Science and Systems, vol. 8, art. 7, 2020. [5] J. W. Smith, J. E. Everhart, W. C. Dickson, W. C. Knowler, and R. S. Johannes, “Using the ADAP learning algorithm to forecast the onset of diabetes mellitus,” in Proc. Annual Symposium on Computer Application in Medical Care, 1988, pp. 261–265. [6] B. Van Calster, D. J. McLernon, M. van Smeden, L. Wynants, and E. W. Steyerberg, “Calibration: the Achilles heel of predictive analytics,” BMC Medicine, vol. 17, art. 230, 2019. [7] E. W. Steyerberg et al., “Assessing the performance of prediction models: a framework for traditional and novel measures,” Epidemiology, vol. 21, no. 1, pp. 128–138, 2010. [8] G. S. Collins, J. B. Reitsma, D. G. Altman, and K. G. M. Moons, “Transparent reporting of a multivariable prediction model for individ-ual prognosis or diagnosis (TRIPOD): the TRIPOD statement,” BMJ, vol. 350, art. g7594, 2015. [9] G. W. Brier, “Verification of forecasts expressed in terms of probability,” Monthly Weather Review, vol. 78, no. 1, pp. 1–3, 1950. [10] A. Niculescu-Mizil and R. Caruana, “Predicting good probabilities with supervised learning,” in Proc. 22nd International Conference on Machine Learning (ICML), 2005, pp. 625–632. [11] J. C. Platt, “Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods,” in Advances in Large Margin Classifiers, A. J. Smola et al., Eds. Cambridge, MA, USA: MIT Press, 1999, pp. 61–74. [12] B. Zadrozny and C. Elkan, “Transforming classifier scores into accurate multiclass probability estimates,” in Proc. 8th ACM SIGKDD Interna-tional Conference on Knowledge Discovery and Data Mining, 2002, pp. 694–699 [13] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th International Conference on Machine Learning (ICML), 2017, pp. 1321–1330. [14] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 4765–4774. [15] S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees,” Nature Machine Intelligence, vol. 2, no. 1, pp. 56–67, 2020. [16] L. S. Shapley, “A value for n-person games,” in Contributions to the Theory of Games II, H. W. Kuhn and A. W. Tucker, Eds. Princeton, NJ, USA: Princeton University Press, 1953, pp. 307–317. [17] E. S?trumbelj and I. Kononenko, “Explaining prediction models and individual predictions with feature contributions,” Knowledge and In-formation Systems, vol. 41, no. 3, pp. 647–665, 2014. [18] M. T. Ribeiro, S. Singh, and C. Guestrin, “‘Why should I trust you?’ Explaining the predictions of any classifier,” in Proc. 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 1135–1144. [19] C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence, vol. 1, no. 5, pp. 206–215, 2019. [20] K. Aas, M. Jullum, and A. Løland, “Explaining individual predictions when features are dependent: more accurate approximations to Shapley values,” Artificial Intelligence, vol. 298, art. 103502, 2021. [21] M. Sundararajan and A. Najmi, “The many Shapley values for model ex-planation,” in Proc. 37th International Conference on Machine Learning (ICML), 2020, pp. 9269–9278. [22] D. Slack, S. Hilgard, E. Jia, S. Singh, and H. Lakkaraju, “Fooling LIME and SHAP: adversarial attacks on post hoc explanation methods,” in Proc. AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2020, pp. 180–186. [23] A. Ghorbani, A. Abid, and J. Zou, “Interpretation of neural networks is fragile,” in Proc. AAAI Conference on Artificial Intelligence, vol. 33, 2019, pp. 3681–3688. [24] D. Alvarez-Melis and T. S. Jaakkola, “On the robustness of interpretabil-ity methods,” in Proc. ICML Workshop on Human Interpretability in Machine Learning (WHI), 2018. [25] A. Kalousis, J. Prados, and M. Hilario, “Stability of feature selection algorithms: a study on high-dimensional spaces,” Knowledge and Infor-mation Systems, vol. 12, no. 1, pp. 95–116, 2007. [26] S. Nogueira, K. Sechidis, and G. Brown, “On the stability of feature selection algorithms,” Journal of Machine Learning Research, vol. 18, no. 174, pp. 1–54, 2018. [27] C. Molnar, Interpretable Machine Learning: A Guide for Making Black Box Models Explainable, 2nd ed., 2022. [28] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5– 32, 2001. [29] T. Chen and C. Guestrin, “XGBoost: a scalable tree boosting system,” in Proc. 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 785–794. [30] F. Pedregosa et al., “Scikit-learn: machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
Copyright © 2026 Ravikumar V, Dr. S. Sasirekha. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84701
Publish Date : 2026-08-24
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
